Remembering and knowing, following Tulving (1985b) and his other early work in which he distinguished between states of consciousness associated with episodic and semantic memory, have been used over the years across an enormous number of studies to capture various phenomenological states. These studies have spanned a variety of areas of study and fields outside traditional memory research, including areas such as neuropsychology, behavioral neuroscience, and the effects of drugs, such as lorazepam or alcohol, on basic memory processes (e.g., Curran, Gardiner, Java, & Allen, 1993; Henson, Rugg, Shallice, Josephs, & Dolan, 1999; Khoe, Kroll, Yonelinas, Dobbins, & Knight, 2000; Moscovitch & McAndrews, 2002). On the basis of an analysis of almost 900 sources that used the R/K paradigm, most researchers have been using it to assess the phenomenological experiences associated with recollection and familiarity (see also Renoult et al., 2019). However, lay participants do not seem to consider these phenomenological experiences of retrieval when defining what it means to remember and to know in contrast to experts in the field of psychology. Instead, lay participants and different types of psychology experts alike associate remembering with the retrieval of event-related information and knowing with retrieval from the knowledge base or semantic memory as Tulving (1985b) originally proposed. Yet very few studies have used the task to discriminate between retrieval from episodic versus semantic memory. Among the most interesting findings from our literature review is the fact that a few studies explicitly avoided using the terms remember and know, choosing instead more neutral terms such as Type A and Type B or opting for labels more transparently descriptive of the construct under investigation (e.g., recollect, familiar). Thus, among users of the R/K paradigm, there is some awareness that the labels might be lacking in face validity, at least when it comes to measuring recollection and familiarity. Study 2 confirms those concerns, indicating that the very labels used in the standard form of the R/K task might be contributing to the variability in how effective the paradigm is in producing clear and replicable assessments of recollection and familiarity (e.g., McCabe & Geraci, 2009; Williams & Lindsay, 2019). The reliance on the specific terms remember and know to capture recollection and familiarity might be problematic given the strong associations between remembering and event memory and knowing and semantic memory, as well as the fact that knowing is associated with higher levels of accuracy, confidence, mastery, and experience-driven learning than remembering. We discuss the current findings as well as their theoretical and more practical implications below, ultimately making some suggestions for the use of the R/K paradigm in future research. Recollection and familiarity As mentioned above, Study 1 indicated that most researchers (95% of the published work examined here) who have used the R/K paradigm have done so to assess the experience of recollection versus that of familiarity. In Study 2, consistent with that finding, experts, especially memory experts, referenced these dimensions more frequently than lay participants. Overall, this may be unsurprising because memory experts have been exposed to, and maybe have even contributed to, the literature on the R/K paradigm that overwhelmingly associates R/K with recollection and familiarity. Although the lay participants showed the same general pattern as other groups (i.e., remembering was more associated with recollection), overall their explanations included very little related to either recollection or familiarity. At the very least, this finding indicates that participants must set aside their common understanding of the terms remember and know when in typical studies that use the R/K paradigm. Because of “the curse of knowledge,” experts may find these terms intuitive to understand, flexibly mapping them onto constructs of interest, and therefore may struggle to understand that participants do not readily do so (Nickerson, 1999). Note that we make no claims about the processes of recollection and familiarity themselves. Rather, these data compel us to echo other researchers’ concerns that “using remember judgments to measure recollection and know judgments to measure familiarity is a crude approach to measuring recollection and familiarity processes relative to more objective methods” (McCabe et al., 2011, p. 1632; see also Wais et al., 2008; Williams & Lindsay, 2019). Event memory, semantic memory, and episodic memory In contrast to the relative failure of R/K to tap recollection and familiarity across participants and experts, there is consensus in defining remembering versus knowing when it comes to event and semantic memory. All groups consistently associated remembering almost exclusively with the retrieval of events and knowing almost exclusively with retrieval from semantic memory. These data and others (Mickes et al., 2013) corroborate Tulving’s original conceptions of the terms (Tulving, 1985b). Intuitively, whether they are memory experts or laypeople examining their own experiences for the first time, people define remembering as the retrieval of experiences of events (and perhaps even recollectionrelated ideas). Knowing, on the other hand, arises from our knowledge base and established long-term storage of information. Note again that we discuss event memory here rather than episodic memory because event memory (Rubin & Umanath, 2015) more broadly encompasses any retrieval related to an event, whereas episodic memory’s ultimate definition is more specified and in line with recollection. Only experts, and mostly memory experts, provided responses that would be considered “episodic” according to strict criteria that would require reference to all of the characteristics of episodic memory (e.g., a unique event that occurred once including the self that one voluntarily retrieves from memory and is accompanied by the sense of reliving); the current data show this pattern in the data regarding recollection. Instead, explanations of “I remember” more frequently referred to thinking back on an event more generally, consistent with event memory (i.e., Rubin & Umanath, 2015). When these data were directly compared in a 2 (Dimension: recollection, event) × 4 (Group: laypeople, memory experts, other cognitive experts, other psychology experts) mixed ANOVA on the inclusion of information relevant to exclusively recollection versus event memory when answering what it means to say “I remember,” the results showed a significant effect of dimension, F(3, 237) = 8.13, MSE = .17, ηp 2 = .03, and of group, F(3, 237) = 6.43, MSE = .24, ηp 2 = .08, but no interaction (p = .39). All participants included more event-related content in their definitions of “I remember” (M = .39) than content related to recollection (M = .28), and more than double the proportion of memory experts included content related to these dimensions compared with laypeople (.51 vs. .21). Thus, the spontaneous use of R/K seems to tap the phenomenological experience of retrieval from different memory stores. In addition, it may be fairer to say that “I remember” taps the retrieval of memories for events more generally than the more specified episodic memory, defined here by recollection, consistent with how Rubin and Umanath (2015) reconceptualized explicit memory. Of course, the open-ended and purposefully vague nature of the question posed to participants and the inherent difference in specificity between recollection and event memory likely influenced their responses to be more general. Knowing The results broadly indicated that participants had more to say about what it means to remember than what it means to know. This pattern was especially strong in experts and more so for memory experts than others. Certainly, for experts, this is in line with the published literature. That is, there is a great deal of work on defining, characterizing, and explaining the phenomenological experiences of recollection, of episodic memory, and so on. Therefore, it is not surprising that these participants more frequently referenced recollection and event memory. What is reflected here is that experts then fail to realize that participants do not have the same level of prior learning regarding memory and its field-specific terminology (e.g., Nickerson, 1999). Often knowledge, semantic memory, and even familiarity are simply considered to be that which is not remembering, a lack of the characteristics that go with recollection, episodic memory, and so on. Even within Tulving’s own conception, knowing transformed from retrieval from the knowledge base (Tulving, 1972, 1984, 1985b, 1987) to memory on “some other basis” than remembering within the context of a recognition task (Gardiner & Java, 1990; Tulving, 1985b, p. 8; Rajaram, 1993; Tulving, 2002). Strack and Förster (1995) observed that the instructions provided to participants regarding when to assign the know judgment often include contradicting examples—one defined as a lack of rememberingrelated mental experiences (e.g., meeting someone on the street and not remembering the exact circumstances under which one first met the person; Rajaram, 1993) and the other defined as retrieval from semantic memory (e.g., that of one’s own name; Gardiner, 1988; Gardiner & Java, 1993). In fact, within the same article, Tulving (1987) goes from discussing semantic memory as general knowledge about the world to the idea that knowing in an episodic memory task—simply a lack of recollection—is retrieval from semantic memory. Thus, knowing could reflect either retrieval from the knowledge base or the absence of recollection in an episodic task. These two flavors of knowing seem fundamentally quite different (Gardiner & Parkin, 1990; Gardiner et al., 1997; Strack & Förster, 1995), and researchers have acknowledged that knowing has been more problematic from the beginning (Gardiner, 1988; Gardiner & Java, 1993; for tables of different usages for know, see Williams & Moulin, 2015; Williams & Lindsay, 2019). For example, to address the possibility that some “know” responses were low-confidence guesses, a “guess” option was introduced early on to separate it from knowing (Dewhurst & Farrand, 2004; Gardiner & Java, 1993; Gardiner et al., 1997; Geraci & McCabe, 2006). Others have introduced “just know” to capture high-confidence knowing (e.g., Conway et al., 1997). Still others have started to use “familiar” instead of “know” (e.g., Parks, 2007) or in combination with guessing (e.g., Bastin et al., 2004) and other options. However, as noted in our review of the literature, the latter modifications are still the exception rather than the norm in the use of the paradigm. Much of the confusion likely arises because of the episodic-recognition task within which R/K is most typically used. Overall, most studies use R/K to examine recollection and familiarity, and very few, outside of those studies on autobiographical memory, do not involve an encoding phase followed by an episodicretrieval phase. What does “I know” mean in this context? Several researchers have raised this very concern (e.g., Barber et al., 2008; Conway et al., 1997; Mickes et al., 2013). Because the paradigm is typically used to assess qualitative or phenomenological aspects of episodic retrieval, framing the use of know as reflecting retrieval from semantic memory appears to be at odds with the constructs under examination. The original conception for knowing as retrieval from one’s knowledge base simply does not make sense (see Rajaram, 1993) in the context of an episodically constrained task, so it is no surprise that other interpretations (e.g., knowing as tapping implicit memory or as familiarity), both on the part of researchers and likely on the part of participants, developed. How can participants be drawing on their knowledge base or semantic memory to explain why they think an item was previously studied in an earlier phase of an experiment? Thus, even if researchers claim to be using the R/K judgments as tapping retrieval from event versus semantic memory, such an application of it cannot be effective in such a task: The task is an event-related one. Whether the participant has been exposed to the word before the study or can define the word—tasks that reflect the engagement of semantic memory—is of no interest. So, knowing as it is defined in natural-language use, as seen in Study 2, is not relevant. Critically, the current data indicate that the memory experience associated with saying “I know” can be defined as much more than just a lack of remembering. Participants here did not define knowing as a lack of recollection or a lack of retrieval of an event. Instead, knowing was defined not only by retrieval from semantic memory, as discussed earlier, but also by constructs related to characteristics of the knowledge base but outside of the traditional dual-processes we examined: confidence, accuracy, mastery, experience, and fluency. Prior work typically associates high confidence with remembering (e.g., Tulving, 1985b; see also Selmeczy & Dobbins, 2014) but sometimes also with knowing (e.g., Conway et al., 1997). In addition, the literature links the ease of processing or retrieval both with automatic processes (familiarity and, thereby, knowing; Rajaram & Geraci, 2000) and with remembering (Algarabel et al., 2003; Kelley & Jacoby, 1998). Here, participants tended to indicate that knowing was associated with “really knowing” something—in other words, they associated knowing (in contrast to remembering) with high confidence, belief in the accuracy of the content, and mastery of a topic, as well as ease of access. These ideas are consistent with some definitions and interpretations of “just know” judgments (Barber et al., 2008; in contrast, see McCabe et al., 2011) and suggests that the prior work done on how participants use remember and know is indeed strongly influenced by the context of an episodic-recognition task. Broad implications The implications of the current work are broad and far-reaching. To quote a reviewer of an earlier version of this work: “Our introductory research methods classes have taught us that reliability, although necessary, is not tantamount to validity and [this work] is really concerned with the validity of the distinction in the eyes of participants.” In this section, we briefly outline some special cases in which these implications warrant further and serious thought. Populations. Older adult participants (typically 65 years or older) have several decades more experience with language than younger adults. Therefore, they might struggle more than younger adults at adapting to using highly familiar terms in a manner that is not consistent with their experience. Furthermore, if overriding a lifelong understanding of what knowing means requires additional cognitive resources, older adults in particular might be at a disadvantage because of documented decreases in cognitive control and inhibitory processes (Park, 2000). Thus, results demonstrating an increased reliance on familiarity-driven responding in aging might be inflated by using measures that tax older adults’ degraded controlled processes. Although we note that converging measures such as the PDP provide consistent evidence for increased familiarity-based responding, that paradigm is also potentially demanding in terms of cognitive resources. The potential implications of the cognitive load imposed by such behavioral measures remain to be determined. Our goal here is to highlight this concern and acknowledge that some aspects of our understanding of basic memory processes in aging are likely shaped by the tools used to examine them. In the almost 80 sources included in our analyses, less than a quarter administered additional tasks to assess RF; this leaves open the question of the extent to which the conclusions in the literature are potentially affected by factors such as cognitive load or task difficulty. Whether the conclusions regarding reliance on familiarity among older adults might change when different paradigms are used remains unclear. For example, only one article in our set used Type A/Type B labels with older adult participants; however, the authors of this article did not compare this label to the traditional R/K labels. There is clearly a need to extend this work to an aging sample and to examine the effects of other labels in this population (cf. Williams & Lindsay, 2019). A second area of particular concern for understanding the underlying constructs under examination when using R/K is evident when working with special populations, as noted by Aggleton et al. (2005): “The first concerns the difficulty that some amnesics may have in subjectively appreciating the difference between ‘remember’ and ‘know’ (Baddeley, Vargha-Khadem, & Mishkin, 2001), associated with the problem of maintaining this difference over a test session” (p. 1821; with regard to cognitive declines, see Bowler, Gardiner, & Grice, 2000; Williams & Moulin, 2015; with regard to healthy aging, see McCabe & Geraci, 2009). This concern highlights that even context may not be enough to support some groups’ ability to understand and use R/K as researchers may want. Language. The potential confusion inherent in relying on the terms remember and know might also be critical when using the paradigm in languages other than English (e.g., languages such as French and Italian have more than one word for knowing; see also McCabe & Geraci, 2009). Among the sources initially identified in Study 1, several (approximately 20) were in languages other than English (this estimate might be conservative if journals in other languages are not indexed in Scopus or Google Scholar). Although a relatively small number, the use of the paradigm in other languages does raise potential questions about what knowing in particular means when multiple words capture subtle differences between forms of knowledge. If some languages distinguish between knowing as retrieval from the knowledge base or semantic memory (e.g., sapere in Italian) and familiarity with someone or something (e.g., conoscere in Italian), this suggests that, conceptually, there are multiple dimensions of knowledge that a single term might struggle to convey. This is clearly an avenue for future research. Cognitive neuroscience research. The current work has critical consequences for neuroscientific research because there is a tendency to use R/K responses as if they provide direct access to the cognitive constructs under investigation rather than treating them with caution as the phenomenological and subjective self-reports that they are. A number of studies and meta-analyses have attempted to identify the cortical and subcortical regions associated with recollection and familiarity (e.g., Cabeza, Ciaramelli, Olson, & Moscovitch, 2008; Eldridge, Knowlton, Furmanski, Bookheimer, & Engel, 2000; Henson et al., 1999; Spaniol et al., 2009; Wais, 2008). The identification of specific structures (e.g., hippocampus or parahippocampal regions) supporting recollection or familiarity is critical for understanding the biological and neurological structures involved in memory performance. However, as others (e.g., Wais, 2008; Wixted & Squire, 2011) have noted, the attribution of recollection to hippocampal regions and of familiarity to surrounding regions depends on a number of factors, such as the measures being used (e.g., source memory tests, confidence ratings, R/K paradigm) or the specific model being tested—for example, high-threshold/dual-process models (Yonelinas, Kroll, Dobbins, Lazzara, & Knight, 1998) versus strength-based models. If R/K responses are confounded with confidence— “remember” responses being generally high-confidence reflections of recollection and “know” responses varying from low- to high-confidence reflections of familiarity—then studies using the R/K paradigm might be capturing differences in confidence-of-recognition judgments (Migo et al., 2012; Wixted & Squire, 2011) rather than the desired constructs of recollection and familiarity. What the results of Study 2 show is that in natural language use, knowing is indeed associated with high levels of confidence. This suggests that in some cases participants’ responses in the R/K paradigm might not be capturing the distinction intended by researchers, adding to the concerns about confounds between confidence, memory strength, and remembering and knowing. Additional concerns that have clear implications for neuroscientific research are that, whereas recollection and familiarity are often considered nonoverlapping constructs in which one process supports a memory decision when the other fails, the results from the R/K paradigm are not always that clear. For example, Eldridge, Engel, Zeineh, Bookheimer, and Knowlton (2005) reported that despite being lower than for “remember” responses, participants were above chance on correctly identifying specific details, such as color and location of a stimulus, for “know” responses, although such characteristics are typically the hallmark of “remember” responses (see also Perfect et al., 1996). Migo et al. (2012) also discuss related issues, such as noncriterial recollection and unconscious recollection. For example, it is possible that a “know” response might include recollected details, such as thoughts one had during encoding, but if the task specifically requires the retrieval of source information, such recollections might not result in a “remember” response or a correct source judgment, leading to an underestimation of recollection. Thus, any measure used to discriminate between recollection and familiarity needs to account for such potential problems. Given that the financialand time-intensive neuroscientific work discussed above depends on the behavioral R/K task successfully and precisely distinguishing the underlying processes, this is a critical issue. Refining the tools being used to measure the constructs of interest is imperative, and these concerns are also relevant to all fields using the paradigm. Finally, as indicated in Study 1, relatively few studies supplement the R/K paradigm with additional measures of recollection and familiarity. This seems to indicate at the very least an implicit assumption on the part of researchers that the paradigm is accurately capturing the underlying constructs and processes. Future use of the R/K paradigm Methodological considerations. A clear conclusion from the current work is that the use of R/K within traditional episodic-recognition tasks can be appropriate if researchers use labels other than remember and know. Some authors have noted that the use of the terms remember and know may sometimes hinder understanding as participants already have a strong idea of what these words mean from outside of the experimental context. Knowing often indicates a high level of confidence in memory and remembering is used in a very broad sense in everyday life. (Migo, Montaldi, Norman, Quamme, & Mayes, 2009, pp. 1446–1447) Others have noted that although the latter terms have been widely used, they could be misinterpreted by participants. The word “remember” in everyday use could denote confidence in one’s memory judgements (irrespective of the presence of recollection) whereas “know”—as per Tulving’s (1985b) original intention, is better suited to testing existing knowledge (i.e., semantic memory) than conveying a sense of familiarity due to an item’s recent exposure. (Tsivilis et al., 2015, p. 6) The current work lends credence to these concerns and corroborates some of the observations. Importantly, we have presented strong evidence to indicate that using the terms as they are commonly used (to address recollection and familiarity) is not intuitive and not how laypeople naturally use the terms. It may be considered a limitation of the current work that we did not specify a “context” in asking participants to define remember and know, whereas there is a great deal of contextualization when instructions are given in the R/K paradigm. However, again, prior work and the compilation of efforts to adjust the paradigm documented in Study 1 show that participants struggle even with context. In addition, we discuss the problem of defining knowing in this standard context above. Moreover, participants may be able to adjust to and use R/K appropriately, but how successfully they do so, whether they maintain the relevant definitions across the duration of the task, and how much of a cognitive load this may add is as yet unknown. The current data are consistent with McCabe and Geraci’s call to use neutral terms such as Type A and Type B experiences instead of remember and know (McCabe and Geraci, 2009; for a critique, however, see Williams & Moulin, 2015) or use terms such as recollect and familiar to capture those phenomenological experiences. Migo et al. (2012) argued that, although the R/K paradigm is at present the recommended way of assessing recollection and familiarity, there is a need for greater consistency and transparency in how it is administered. They summarized five key elements that need to be included in work using the R/K paradigm: verbatim instructions used that clearly state what the terms remember and know (or alternatives) mean; how participants’ comprehension of the instructions was assessed; how compliance with the instructions was assessed; whether any participants’ data were omitted and why; and how familiarity was computed. Our review of the literature confirms and highlights these concerns. Specifically, there is a large degree of variability in the instructions used and the details provided, and that variability actually affects how participants assign the terms to their phenomenological experiences (Williams & Lindsay, 2019). Adherence to providing similar methodological details would be a good start. One simple possible solution, as suggested by Migo et al. (2012), would be to ask participants at the end of the study what they meant when they used each term and exclude those who do not show full understanding from the analyses. Very few studies using the R/K paradigm thus far have included a posttest assessment (< 10%), and it is unclear which of these few use the posttest assessment as an exclusion criterion. In contrast, in studies regarding prospective memory, it is common practice to ask participants at the completion of the study what the assigned prospective memory tasks were to ensure that they hold that intention in memory across the duration of the study (e.g., McDaniel, Shelton, Breneiser, Moynan, & Balota, 2011). Participants who fail to recall the prospective task are then often not included in the analyses (e.g., Kvavilashvili, Kornbrot, Mash, Cockburn, & Milne, 2009; McDaniel et al., 2011). The rationale in that area is that assessing prospective memory and intentions versus success in performing future actions depends critically on participants actually understanding and maintaining those intentions, neither of which is trivial to the research at hand (Einstein, McDaniel, Williford, Pagan, & Dismukes, 2003). For R/K, if participants provide explanations for remembering and knowing that align with recollection and familiarity (perhaps according to the coding scheme included in the current work), their data can be included in the analyses. If not, their data should be removed from the study because their responses in the task may not reflect what the researchers are hoping to assess, degrading the validity of the work. Thus, given the potential lack of clarity and need for extensive instruction inherent to the R/K paradigm, it appears that the inclusion of a simple posttest assessment to ensure participants were consistently applying the terms remember and know in the ways intended by the experimenters would increase the confidence in the validity of the measure. Other similar solutions include using a form of catch trials wherein participants are asked to justify random responses as a check (Gardiner et al., 1997). Rotello et al. (2005) suggest in a footnote that perhaps only studies with very low false-alarm rates for “remember” responses should be used because such a criterion would ensure that participants understood the instructions properly. Remarkably, they then point out that this would disqualify approximately a third of all R/K work in the literature.2 Thus, Geraci et al. (2009) observed that almost a fifth of their participants did not understand the instructions when asked after the study. A more intensive approach may be to eschew relying on self-assessment of R/K altogether and simply ask for verbal explanations of “old” judgments with the assessment of the underlying processes decided by researchers, as McCabe et al. (2011) and Selmeczy and Dobbins (2014) have done. Conceptual considerations. We begin this section with a quote from Tulving & Nilsson (1979) as they concluded an article entitled “Memory Research: What Progress?”: Finally, we should be willing—perhaps “have the courage” would be a more appropriate expression— to reject ideas and hypotheses that are at variance with the data. Instead, frequently the hypotheses incompatible with the data are maintained or just mended, and mended again when they encounter further difficulties. Mending usually takes the form of adding an additional wrinkle, another qualification, or another parameter or two. If such recurrent modification continues for a while, the explanation may eventually collapse under its own weight; but in the meantime its existence has stood in the way of an active search for a better one. (p. 31) Given the extensive methodological “mending” suggestions above and in prior work, does R/K need to be discarded completely? In fact, we would argue that no, the paradigm as it currently exists does not need to be discarded. However, at the very least, its use does need to be carefully constrained in the future (see also Williams & Lindsay, 2019). Above, we addressed changes to the methodological implementation of the paradigm if researchers intend to use it to measure RF. Now, perhaps more critically, we discuss its conceptual underpinnings. How can R/K be used validly and effectively? Our suggestions can be summarized as follows: Use R/K (a) to better understand the phenomenological experiences of autobiographical memory and (b) to examine the transition of memories from event memory to the knowledge base. These terms can be meaningful and used effectively in the context of retrieval of autobiographical memories wherein one can remember, recollect, and mentally relive past events versus simply know that they occurred (for a review, see Moulin et al., 2013). Likewise, Picard et al. (2013) implemented the terms in the context of a very rich experiential task in which the terms’ meanings were more consistent with retrieval of events- versus knowledge-based experiences. Still, caution is warranted, because in other work on autobiographical memory, remember judgments are more highly correlated with belief in the accuracy of the memory compared with experiences of reliving (Rubin, Schrauf, & Greenberg, 2003; Rubin & Siegler, 2004). R/K can also be used when the task is meant to examine the contents of the knowledge base (e.g., Barber et al., 2008). For example, Conway et al. (1997) examined the longer-term process of learning across time and the transformation of new learning from being linked to event-related associations to knowledgerelated associations, using just knowing to capture semantic memory or the knowledge base versus familiar to tap low-confidence, weak memory traces (along with remember and guess; see also Barber et al., 2008). Furthermore, it is clear that remembering and knowing do naturally reflect distinct phenomenological states: that of retrieval of events versus that of retrieval of knowledge. Thus, our fundamental claim is that the use of these terms can be a valuable tool for exploring the phenomenology of retrieval as long as the use of the terms is understood and agreed on by both participants and researchers (Bahrick, Baker, Hall, & Abrams, 2011; Coane & Umanath, 2019).